What does it mean to learn?
When you were learning to ride a bike, nobody wrote you a manual with instructions like "apply 2.3 Newtons of force to the left pedal at exactly 38 degrees." That would have been useless. Instead, you got on the bike, tried, fell, adjusted, tried again, and slowly your body figured out the balance. Trial, error, adjustment, repeat.
Machine learning works the same way. You do not program rules into the model. You give it data, let it make predictions, show it how wrong those predictions were, and let it adjust. Do that thousands or millions of times and the model gets good at the task. The machine is learning from feedback, exactly like you learned to ride that bike.
The difference is scale. A human might need a few hundred attempts to learn a skill. A machine can run through millions of examples in minutes.
"Machine learning is a field of study that gives computers the ability to learn without being explicitly programmed."
Arthur Samuel, 1959 — the person who coined the term "machine learning"The learning loop
Every machine learning model, from the simplest linear regression to the largest language model, follows the same basic cycle. Understanding this loop is the most important conceptual foundation in this entire phase.
That cycle runs over and over, adjusting the model slightly each time, until the predictions are good enough. The number of times you run the full loop is called an epoch. Running for 10 epochs means the model has seen the full training dataset 10 times.
What the model is actually adjusting
Inside every machine learning model there are numbers called parameters (also called weights and biases). These are the knobs the model turns during training. At the start, they are set randomly. By the end of training, they hold the learned pattern.
Think of a linear model predicting house prices. Its prediction might look like this:
price = (w1 × size) + (w2 × bedrooms) + (w3 × location) + b
# w1, w2, w3 are the weights (parameters the model learns)
# b is the bias (another learned parameter)
# At the start: w1=0.1, w2=0.3, w3=-0.2, b=50
# After training: w1=180, w2=25000, w3=42000, b=-15000The model starts with random values for all those weights. It is terrible at first. After training, those weights encode the actual relationship between features and price that the model discovered in the data. Training is nothing more than finding the right values for these numbers.
How does it know which way to adjust?
This is where gradient descent comes in. Gradient descent is the algorithm that figures out how to adjust the weights to reduce the error. The intuition is simple, even if the full mathematics runs deeper.
Imagine the curve as a valley. The ball starts somewhere random on the slope. Gradient descent figures out which way is downhill and takes a small step in that direction. Repeat enough times and the ball settles at the lowest point, where the error is smallest. That lowest point is your trained model.
The size of each step is called the learning rate. Too large and you overshoot the bottom and bounce around. Too small and training takes forever. Choosing a good learning rate is one of the most important practical decisions in machine learning.
Parameters vs hyperparameters
There are two kinds of numbers in a machine learning model and it is worth being clear about the difference because the terms come up constantly.
Parameters are what the model learns. Hyperparameters are what you decide before learning begins. A chef does not choose how the dish tastes while eating it. They choose the recipe, the oven temperature and the cooking time before the dish goes in. Those choices are the hyperparameters. What comes out is shaped by the parameters the process discovers.
Your first model in Scikit-learn
All of this theory becomes much more concrete once you see the loop in code. Scikit-learn makes this remarkably clean. Here is a full machine learning pipeline in about 10 lines.
from sklearn.linear_model import LinearRegression from sklearn.model_selection import train_test_split from sklearn.metrics import mean_squared_error import numpy as np # 1. Prepare your features and label X = df[['size', 'bedrooms', 'location_score']] y = df['price'] # 2. Split into training and test sets X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42) # 3. Create the model (a hyperparameter choice: which model?) model = LinearRegression() # 4. Train: the loop runs inside here model.fit(X_train, y_train) # 5. Predict on unseen data predictions = model.predict(X_test) # 6. Measure the error error = mean_squared_error(y_test, predictions) print(f"Mean squared error: {error:.2f}")
Notice that you never write the learning loop yourself. model.fit() runs the entire gradient descent process for you. Scikit-learn handles all the maths internally. Your job is to prepare clean data, choose a sensible model, and then evaluate whether the result is good enough.
Calling model.fit() is like hiring an experienced chef and handing them a recipe book full of examples. The chef studies all the examples, figures out the patterns, and develops their own internal skills. You do not watch every step of that process. When they are done, you call model.predict() and they cook you a meal based on everything they learned.
What training cannot do
It is worth being honest about the limits of this process. The learning loop optimises the model to perform well on its training data. That is not the same as understanding the world. A model that has learned from historical house prices knows nothing about what a house physically is. It knows which numbers tend to go together.
This is why a model trained on data from one city can perform poorly in another. It has learned the patterns in what it saw, not the underlying rules of reality. Keeping this distinction in mind will save you from over-trusting your models and will help you diagnose problems when they arise.